Papers with dialogue metrics
Explaining Dialogue Evaluation Metrics using Adversarial Behavioral Analysis (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing frameworks for dialogue model evaluation are lacking to investigate these biases . a number of dialogue metrics are biased and can cause unforeseen problems . |
| Approach: | They propose an adversarial test-suite which generates problematic variations of various dialogue aspects using automatic heuristics. |
| Outcome: | The proposed test-suite generates problematic variations of various dialogue aspects using automatic heuristics. |
Exploring the Impact of Human Evaluator Group on Chat-Oriented Dialogue Evaluation (2024.lrec-main)
Copied to clipboard
| Challenge: | Evaluator groups such as domain experts, university students, and crowdworkers have been used to assess and compare chat-oriented dialogue systems. |
| Approach: | They analyze the impact of evaluator groups on dialogue system evaluation by testing 4 state-of-the-art dialogue systems using 4 distinct evaluer groups. |
| Outcome: | The proposed evaluations show that the evaluator group impact is not seen for Pairwise, and that it is beneficial for certain metrics. |
IM^2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue Evaluation (2022.emnlp-main)
Copied to clipboard
| Challenge: | Evaluation metrics for dialogue systems are expensive and time-consuming . current evaluation metrics focus on a single quality or several qualities . |
| Approach: | They propose an interpretable, multi-faceted, and controllable framework to combine dialogue metrics which are good at measuring different qualities. |
| Outcome: | The proposed framework integrates a large number of evaluation metrics to improve the performance of the model. |